feat(srt): check srt-slurm recipes against master images before sweep dispatch / feat(srt):在派发 sweep 前检查 srt-slurm 配方与主配置镜像是否一致 - #3624
chunfangamd wants to merge 2 commits into
Conversation
| if [ -f "$PRIORITY_ROOT/infx/srt_slurm/preflight.py" ] && [ -f "$MEASURED_ROOT/infx/srt_slurm/preflight.py" ]; then | ||
| git -C "$MEASURED_ROOT" submodule update --init utils/srt-slurm | ||
| CONFIG_JSON=$(printf '%s' "$CONFIG_JSON" | env PYTHONPATH="$PRIORITY_ROOT:$MEASURED_ROOT/utils/srt-slurm/src" \ | ||
| uv run --no-project --exclude-newer PT12H --python 3.12 --with pydantic --with pyyaml \ | ||
| --with marshmallow --with marshmallow-dataclass --with ruamel.yaml --with requests \ | ||
| python -P -m infx.srt_slurm.preflight --root "$MEASURED_ROOT") |
There was a problem hiding this comment.
🔴 This runs the "trusted" preflight check with srtctl imported from the PR's own utils/srt-slurm submodule checkout, letting a malicious PR execute code inside the trusted-tooling process instead of only the sandboxed measured tree. check_multi_node -> selected_recipes does from srtctl.core.config import generate_override_configs, and PYTHONPATH is $PRIORITY_ROOT:$MEASURED_ROOT/utils/srt-slurm/src (e2e-tests.yml:336) while only $MEASURED_ROOT ever gets git submodule update --init utils/srt-slurm (line 335) — $PRIORITY_ROOT never does, so srtctl can only resolve from the PR-controlled submodule commit. -P doesn't block PYTHONPATH imports. Fix: give PRIORITY_ROOT its own trusted srt-slurm checkout (or vendor/pin the srtctl it imports) so the trusted interpreter never imports a package whose source lives only in the untrusted MEASURED_ROOT tree; same pattern in profile.yml:111-116.
Why this was flagged
The whole point of checking out .ci-priority/PRIORITY_ROOT separately from MEASURED_ROOT in this job is to run decision logic with code the PR cannot modify. This diff defeats that: infx.srt_slurm.preflight (loaded from PRIORITY_ROOT) calls infx.srt_slurm.synthetic_acceptance.selected_recipes, which lazily imports srtctl.core.config.generate_override_configs; that package is only ever present via $MEASURED_ROOT/utils/srt-slurm/src (e2e-tests.yml:335-336), which is the PR's own branch/submodule pointer. A PR that points its utils/srt-slurm submodule gitlink at a malicious commit gets that code executed inside the trusted PRIORITY_ROOT python process when get-jobs runs (same for profile.yml:111-116). python -P only disables automatic sys.path prepending, not PYTHONPATH, so this import is not blocked. The output of this step feeds CONFIG_JSON/job outputs that downstream benchmark-dispatch jobs (which do use secrets.INFERENCEX_OFFICIAL_RO_HF_TOKEN, secrets.MODAL_TOKEN_ID/SECRET) consume, so compromising this step can also poison what those privileged jobs run.
Verification: e2e-tests.yml:335 runs git -C "$MEASURED_ROOT" submodule update --init utils/srt-slurm (PR-controlled submodule), and :336 sets PYTHONPATH="$PRIORITY_ROOT:$MEASURED_ROOT/utils/srt-slurm/src", so srtctl resolves only from the measured tree.
b95b930 to
eddfb6b
Compare
…atch run-sweep, e2e-tests and profile now pipe their matrix through infx.srt_slurm.preflight after benchmark_schema. A single-node point must select exactly one variant through the runtime's select_recipe. Every CONFIG_FILE and EVAL_CONFIG_FILE recipe a multi-node point can launch must select at least one variant, and its worker containers must resolve to the images the launcher stages: the point's image or a cluster alias, and PREFILL_IMAGE for a TileRT prefill role. A TileRT frontend or client must run one of those images, and a benchmark client that names another tag of the point's image is reported. Recipe paths must stay inside their recipe trees. A mismatch fails setup before any canary or benchmark job is dispatched. In e2e-tests and profile, the validator and its srtctl come from the trusted tree, which reads the measured tree only as data, and the check always runs. require_launcher already rejects revisions older than the Python launcher, which are also the revisions whose runner config this tooling cannot read. run-sweep、e2e-tests 和 profile 现在在 benchmark_schema 之后把 matrix 交给 infx.srt_slurm.preflight 检查。单节点点必须通过运行时的 select_recipe 恰好 匹配一个 variant。多节点点可能启动的每个 CONFIG_FILE 和 EVAL_CONFIG_FILE 配方必须至少选中一个 variant,其 worker container 必须解析到 launcher 暂存 的镜像:该点的镜像或 cluster alias;TileRT 的 prefill role 则必须是 PREFILL_IMAGE。TileRT 的 frontend 和客户端必须使用这两个镜像之一;benchmark 客户端如果写了该点镜像的另一个 tag,会被报告。配方路径必须位于各自的配方 目录内。任何不一致都会让 setup 在派发 canary 或 benchmark 任务之前失败。 在 e2e-tests 和 profile 中,validator 及其 srtctl 都来自可信代码树,被测代码 树只作为数据读取,并且检查始终运行。require_launcher 已经拒绝早于 Python launcher 的版本,而本工具无法读取 runner 配置的正是这些版本。 Co-authored-by: Cursor <[email protected]>
… points The srt-slurm launcher (lanes.config_file) launches EVAL_CONFIG_FILE only for an eval-only point that sets it; every other multi-node point needs a non-empty CONFIG_FILE and otherwise fails only on the GPU runner. The preflight now reports such points before dispatch, and still checks every recipe a point references. srt-slurm launcher(lanes.config_file)只会为设置了 EVAL_CONFIG_FILE 的 eval-only 点启动该配方;其他多节点点都需要非空的 CONFIG_FILE,否则要到 GPU runner 上才会失败。preflight 现在会在派发前报告这类点,并继续检查点 所引用的每个配方。 Co-authored-by: Cursor <[email protected]>
eddfb6b to
5930300
Compare
Description
Follow-up to #3567; the two PRs belong together. #3567 aligns the two srt-slurm recipes whose container had drifted from the master-config
image. This PR adds the check that would have caught the drift: each sweep now binds every planned srt-slurm point to the recipe variants it would launch, and stops before dispatching any job if a point doesn't bind.Problem. An srt-slurm point names its images in two places: the master config's
image(plusPREFILL_IMAGEfor TileRT) and the containers in the recipe it launches. An image bump has to edit both, as #3446 did, but nothing compares them before benchmark jobs start.infx/srt_slurm/single_node.pyrejects a mismatch (Single-node SRT image: recipe/matrix ...) only after the benchmark job has started on its cluster runner, prepared an srt-slurm checkout and installed srtctl. [AMD] Update GLM-5.2 MI355X image to 20260924 daily / 更新 GLM-5.2 MI355X 镜像至 20260924 daily #3446's first sweep failed this way. So would every point of the two keys fix(srt): align B200 and MI355X recipe images with master configs / fix(srt):对齐 B200 与 MI355X 配方镜像与主配置 #3567 fixes, because feat(srt): run AgentX on native srt-slurm #3428 ported their recipes from the legacy scripts as they were before the image bumps in config(dsv41flash): re-sweep MI355X on the first ROCm 10.0 nightly carrying vllm#58510 / 在首个包含 vllm#58510 的 ROCm 10.0 nightly 上重新扫描 MI355X #3420 and [Klaud Cold] Update dsv4-fp4-b200-sglang-agentic-hicache-mtp SGLang image to v0.5.20-cu130 / 将 dsv4-fp4-b200-sglang-agentic-hicache-mtp 的 SGLang 镜像更新至 v0.5.20-cu130 #3334.containers:map as a literal image, so a stale recipe benchmarks a different image from the one the master config records.#3567 first added a pytest scan for this. Review found that it couldn't gate
run-sweep, compared image sets per recipe instead of the variant each point selects, skippedEVAL_CONFIG_FILE, and pinned checked-in config against theAGENTS.mdtest rules. #3567 dropped the scan, and this PR replaces it.Changes
infx/srt_slurm/preflight.py(python -m infx.srt_slurm.preflight) reads the matrix on stdin. If every srt-slurm point binds, it writes the matrix unchanged to stdout; otherwise it prints each problem once with the points it affects and exits 1.benchmark-tmpl.ymlexports for the point and calls the runtime's ownselect_recipe, which must match exactly one variant on engine, model, image, precision, parallelism, GPU count, speculative decoding, concurrency and KV offloading. The check runs per point, so a recipe shared by keys with different images is checked against each key's image.CONFIG_FILE, unless it is eval-only and setsEVAL_CONFIG_FILE. That is the recipelanes.config_filelaunches; any other point would fail only on the GPU runner. The validator expands everyCONFIG_FILEandEVAL_CONFIG_FILEa point references with srtctl's own override expansion. A reference that selects no variant is reported, because srtctl would submit nothing. In every variant:model.containerand eachroles.<role>.containermust be the point's image (in either registry spelling) or a container alias thatconfigs/runners.yamldefines for the runner's cluster. A TileRT point must setPREFILL_IMAGE, and its prefill role must be exactly that image, the only prefill name the launcher stages.frontend.container_imageorbenchmark.container_imagemust be the decode image (or an alias) orPREFILL_IMAGE.identity.container.imagemust name the point's image.benchmark.container_imagethat names another tag of the point's image is reported, because srtctl would pull that literal instead of the staged image. An unrelated client image is allowed, and other frontends may pin an image of their own: the ATOM agentic recipes pinrocm/atom-dev@sha256:f00a…for theiratomeshfrontend.srt-recipemust resolve insidebenchmarks/single_node/srt-slurm-recipes, and a multi-node reference must be arecipes/path that resolves insidebenchmarks/multi_node/srt-slurm-recipes, so..and absolute paths are rejected.run-sweep.yml:setupinitializes the srt-slurm submodule and runs the validator afterbenchmark_schema --plan, beforeci_priority. If the validator fails,setupfails, socanary-selectand the benchmark jobs, which all require a successfulsetup, never start.e2e-tests.yml(whichtrusted-external-sweep.ymlandclaude.ymlalso dispatch) andprofile.yml: the validator and its srtctl come from the trusted tooling tree (PRIORITY_ROOTand its own srt-slurm submodule), and the measured tree is read only as data (--root "$MEASURED_ROOT"). The step always runs.require_launcher, earlier in the same step, already rejects revisions older than the Python launcher, and those are also the revisions whoserunners.yamlthis tooling cannot read.infx/tests/srt_slurm/test_preflight.pyhas 36 cases with small temporary recipes and runner inventories, per theAGENTS.mdtest rules. They cover:EVAL_CONFIG_FILE, or onlyCONFIG_FILE;EVAL_CONFIG_FILEmismatch;:baseis named, and a variant withoutmodel.container;PREFILL_IMAGE;..and unprefixed recipe paths, and a missing recipe.A sweep that selects the B200 key without #3567's fix stops at
setupwith:Changes since the first review
e2e-tests.ymlandprofile.ymlno longer import srtctl from the measured tree's submodule, which a PR could point at any code (Claude Code Review and the external review).preflight.py.base.benchmark.container_imagemust follow the point's image unless it names an unrelated client image.CONFIG_FILEis reported unless it is eval-only and setsEVAL_CONFIG_FILE(second review).main1f60c15b5, and the second-review fix is its own bilingual commit.Validation
All results below are on this branch rebased onto
main1f60c15b5.test_preflight.py: 36 passed. Each of 21 deliberate regressions in the validator fails at least one test, including skipping TileRT, treating the prefill role like the others, accepting zero variants, letting a throughput point launchEVAL_CONFIG_FILE, dropping either path check, and reading the runner inventory eagerly.ci.ymlcommands, run locally: Lint is clean, and Tests has 2,380 passed. Twotest_slurm_clicases fail only on the test host because it has a realsacct; they fail the same way on cleanmain.zizmorcommand reports nothing on a clean checkout without submodules, as CI checks it out.dsv4-fp4-b200-sglang-agentic-hicache-mtp(12 points) is fixed by fix(srt): align B200 and MI355X recipe images with master configs / fix(srt):对齐 B200 与 MI355X 配方镜像与主配置 #3567.dsv41flash-fp4-mi355x-vllm-agentic-dsparkis now consistent onmainafter [AMD][MI355X] DSv4.1-Flash vLLM: enable ROCm paged MXFP4 sparse indexer / DSv4.1-Flash vLLM:启用 ROCm 分页 MXFP4 稀疏 indexer #3571.glm5.2-fp8-mi325x-sglang-agentic-mtp(7 points) still drifts onmain.dsv4-fp4-mi355x-sglang-disagg-agentic-umbp-dspark(6 points, one per override): its recipe has run its benchmark client onlmsysorg/sglang-rocm:v0.5.19-rocm720-mi35x-20260907since feat(srt): run AgentX on native srt-slurm #3428, while the master image isv0.5.20-rocm720-mi35x-20260923.run-sweepsetupreplayed in a throwaway worktree with the workflow's exact commands underbash -e:ci_priority, listing 14 points; with fix(srt): align B200 and MI355X recipe images with master configs / fix(srt):对齐 B200 与 MI355X 配方镜像与主配置 #3567's recipe, the matrix passes through byte-identical.glm5.3-fp8-mi355x-tilert-agenticmaster image from0.1.6to0.1.7exits 1 onmodel.container,roles.decode.containerandfrontend.container_image; bumping the recipe too passes.minimaxm3-fp4-gb300-dynamo-vllm-agentic-mtp-disaggpasses: its plan has 6 throughput points and 10eval-onlypoints. With itsCONFIG_FILElines removed andEVAL_CONFIG_FILEkept, the 6 throughput points fail and the eval-only points still pass.e2e-testsandprofile: their preflight blocks, extracted verbatim from the workflows, ran underbash -eagainst a measuredmaincheckout whose srtctl was replaced by a trap that exits on import, and which has nopreflight.py. The trap never fired. The B200 drift was reported, and with fix(srt): align B200 and MI355X recipe images with master configs / fix(srt):对齐 B200 与 MI355X 配方镜像与主配置 #3567's recipe the matrix passed unchanged. With the oldPYTHONPATHthe trap fired.Notes for reviewers
dsv4-fp4-b200-sglang-agentic-hicache-mtpfails atsetup; it would fail at runtime anyway.glm5.2-fp8-mi325x-sglang-agentic-mtpanddsv4-fp4-mi355x-sglang-disagg-agentic-umbp-dsparkwill fail atsetupwhen a sweep selects them, until their recipes are fixed. Fixing the dspark client image changes what the benchmark runs, so it needs its own changelog entry and sweep. Only the points a sweep plans are checked, so neither blocks unrelated PRs.single_node_environmentmirrors howbenchmark-tmpl.ymlexports matrix fields to the job. A change to that mapping needs the same change here.run-sweeponly triggers on PRs that editperf-changelog.yaml, so this PR's own checks don't run the newsetupstep. The replays above stand in for that.perf-changelog.yamlentry, because no benchmark config or recipe changes.@SemiAnalysisAI/core.AI model disclosure
Related Issue
No issue. Follow-up to #3567. Related: #3428, #3446.
Type of Change
Checklist
inferencex-e2e/perf-changelog.yamland have not edited historical entriesOWNER/MEMBER/COLLABORATOR) has commented/use <run_id>(or the legacy/reuse-sweep-run) on this PR. Do this only once there is a final full sweep that is all green with evals passing, since after this comment the sweep label will no longer automatically kick off new sweeps. Remove and re-add the label to force one.中文
改动说明
本 PR 是 #3567 的后续,两者是一个整体。#3567 修正了两个 container 与主配置
image不一致的 srt-slurm 配方;本 PR 加上本可以发现这种不一致的检查:每次 sweep 都会把每个计划中的 srt-slurm 点绑定到它实际会启动的配方 variant,任何一个点绑定失败,就在派发任何任务之前停止。问题: srt-slurm 点的镜像写在两个地方:主配置的
image(TileRT 还有PREFILL_IMAGE),以及它所启动的配方里的各个 container。升级镜像必须同时修改两处(如 #3446),但在 benchmark 任务开始之前,没有任何检查比较这两处。infx/srt_slurm/single_node.py要等 benchmark 任务已在集群 runner 上启动、准备好 srt-slurm checkout 并安装 srtctl 之后,才会拒绝不一致的点(Single-node SRT image: recipe/matrix ...)。[AMD] Update GLM-5.2 MI355X image to 20260924 daily / 更新 GLM-5.2 MI355X 镜像至 20260924 daily #3446 的第一次 sweep 就是这样失败的。fix(srt): align B200 and MI355X recipe images with master configs / fix(srt):对齐 B200 与 MI355X 配方镜像与主配置 #3567 修复的两个 key 的每个点也会这样失败,因为 feat(srt): run AgentX on native srt-slurm #3428 移植这些配方时,依据的是 config(dsv41flash): re-sweep MI355X on the first ROCm 10.0 nightly carrying vllm#58510 / 在首个包含 vllm#58510 的 ROCm 10.0 nightly 上重新扫描 MI355X #3420 和 [Klaud Cold] Update dsv4-fp4-b200-sglang-agentic-hicache-mtp SGLang image to v0.5.20-cu130 / 将 dsv4-fp4-b200-sglang-agentic-hicache-mtp 的 SGLang 镜像更新至 v0.5.20-cu130 #3334 升级镜像之前的旧脚本。containers:映射里的配方 container 当作字面镜像使用,所以过期的配方会测试一个与主配置记录不同的镜像。#3567 最初为此加了一个 pytest 扫描。Review 指出:它无法拦住独立的
run-sweep;它按配方比较镜像集合,而不是每个点实际选中的 variant;它没有检查EVAL_CONFIG_FILE;并且在测试里固定了 checked-in 配置,违反AGENTS.md的测试规则。#3567 已移除该扫描,由本 PR 取代。改动:
infx/srt_slurm/preflight.py(python -m infx.srt_slurm.preflight)从 stdin 读入 matrix。所有 srt-slurm 点都能绑定时,原样输出到 stdout;否则每个问题只打印一次并列出受影响的点,然后以退出码 1 结束。benchmark-tmpl.yml为该点导出的环境变量,调用运行时自己的select_recipe,必须在引擎、模型、镜像、精度、并行配置、GPU 数、投机解码、并发和 KV offloading 上恰好匹配一个 variant。由于逐点检查,被多个不同镜像的 key 共享的配方会分别对照每个 key 的镜像。CONFIG_FILE,除非它是 eval-only 点并设置了EVAL_CONFIG_FILE。这正是lanes.config_file会启动的配方;其他情况要到 GPU runner 上才会失败。validator 用 srtctl 自己的 override 展开逻辑,展开点所引用的每个CONFIG_FILE和EVAL_CONFIG_FILE。展开出 0 个 variant 的引用会被报告,因为 srtctl 不会提交任何任务。在每个 variant 中:model.container和每个roles.<role>.container必须是该点的镜像(两种 registry 写法均可),或configs/runners.yaml为该 runner 所在 cluster 配置的 container alias。TileRT 点必须设置PREFILL_IMAGE,其 prefill role 必须恰好是这个镜像,这是 launcher 为 prefill 暂存的唯一名字。frontend.container_image或benchmark.container_image必须是 decode 镜像(或 alias)或PREFILL_IMAGE。identity.container.image时,它必须是该点的镜像。benchmark.container_image如果写的是该点镜像的另一个 tag,就会被报告,因为 srtctl 会拉取这个字面镜像,而不是使用暂存的镜像。无关的客户端镜像可以使用;其他 frontend 也可以固定自己的镜像,例如 ATOM agentic 配方为atomeshfrontend 固定了rocm/atom-dev@sha256:f00a…。srt-recipe必须解析到benchmarks/single_node/srt-slurm-recipes之内;多节点引用必须是recipes/路径,并解析到benchmarks/multi_node/srt-slurm-recipes之内,因此..和绝对路径都会被拒绝。run-sweep.yml:setup先初始化 srt-slurm submodule,在benchmark_schema --plan之后、ci_priority之前运行 validator。validator 失败时setup失败,canary-select和所有要求setup成功的 benchmark 任务都不会启动。e2e-tests.yml(trusted-external-sweep.yml和claude.yml也通过它派发)和profile.yml:validator 及其 srtctl 都来自可信工具树(PRIORITY_ROOT及其自己的 srt-slurm submodule),被测代码树只作为数据读取(--root "$MEASURED_ROOT")。该步骤始终运行。同一步骤中更早执行的require_launcher已经拒绝早于 Python launcher 的版本,而本工具无法读取runners.yaml的也正是这些版本。infx/tests/srt_slurm/test_preflight.py按AGENTS.md的测试规则,用 36 个用例,只使用临时目录中的小型配方和 runner 配置,覆盖:EVAL_CONFIG_FILE、只有CONFIG_FILE时的情况;EVAL_CONFIG_FILE不一致;:base时展开出 0 个 variant 的配方,以及缺少model.container的 variant;PREFILL_IMAGE的 TileRT 点;..和缺少recipes/前缀的配方路径,以及缺失的配方。上方的示例输出,是一个选中 B200 key、但没有 #3567 修复的 sweep 在
setup停止时打印的内容。相对第一次 review 的改动:
e2e-tests.yml和profile.yml不再从被测代码树的 submodule 导入 srtctl,因为 PR 可以让它指向任意代码(Claude Code Review 和外部 review 都指出了这一点)。preflight.py。base。benchmark.container_image必须跟随该点的镜像,除非它是无关的客户端镜像。CONFIG_FILE的多节点点会被报告,除非它是 eval-only 点并设置了EVAL_CONFIG_FILE(第二次 review)。main的1f60c15b5,第二次 review 的修复是单独的双语 commit。验证:
以下结果都基于 rebase 到
main1f60c15b5之后的本 branch。test_preflight.py:36 个通过。对 validator 故意做的 21 种破坏,每一种都至少让一个测试失败,包括跳过 TileRT、把 prefill role 当作普通 role、接受 0 个 variant、让 throughput 点启动EVAL_CONFIG_FILE、去掉任一路径检查,以及提前读取 runner 配置。ci.yml的命令:Lint 干净,Tests 2,380 个通过。两个test_slurm_cli用例只在测试机上失败,因为它装有真实的sacct;在干净的main上同样失败。zizmor,没有任何报告。dsv4-fp4-b200-sglang-agentic-hicache-mtp(12 个点)由 fix(srt): align B200 and MI355X recipe images with master configs / fix(srt):对齐 B200 与 MI355X 配方镜像与主配置 #3567 修复。dsv41flash-fp4-mi355x-vllm-agentic-dspark在 [AMD][MI355X] DSv4.1-Flash vLLM: enable ROCm paged MXFP4 sparse indexer / DSv4.1-Flash vLLM:启用 ROCm 分页 MXFP4 稀疏 indexer #3571 之后已在main上一致。glm5.2-fp8-mi325x-sglang-agentic-mtp(7 个点)在main上仍不一致。dsv4-fp4-mi355x-sglang-disagg-agentic-umbp-dspark(6 个点,每个 override 一个):它的配方自 feat(srt): run AgentX on native srt-slurm #3428 起就用lmsysorg/sglang-rocm:v0.5.19-rocm720-mi35x-20260907运行 benchmark 客户端,而主配置镜像是v0.5.20-rocm720-mi35x-20260923。bash -e重放run-sweep的setup:ci_priority,列出 14 个点;带上 fix(srt): align B200 and MI355X recipe images with master configs / fix(srt):对齐 B200 与 MI355X 配方镜像与主配置 #3567 的配方后,matrix 逐字节不变地通过。glm5.3-fp8-mi355x-tilert-agentic的主配置镜像从0.1.6升到0.1.7时,以退出码 1 报出model.container、roles.decode.container和frontend.container_image;同时升级配方后通过。minimaxm3-fp4-gb300-dynamo-vllm-agentic-mtp-disagg通过:它的 plan 有 6 个 throughput 点和 10 个eval-only点。删掉它的CONFIG_FILE、只保留EVAL_CONFIG_FILE后,6 个 throughput 点失败,eval-only 点仍然通过。e2e-tests和profile:从 workflow 中原样抽取的 preflight 命令块,在bash -e下针对被测的maincheckout 运行;该 checkout 的 srtctl 被替换成导入即退出的陷阱,并且没有preflight.py。陷阱始终没有触发:报出了 B200 的不一致;带上 fix(srt): align B200 and MI355X recipe images with master configs / fix(srt):对齐 B200 与 MI355X 配方镜像与主配置 #3567 的配方后,matrix 原样通过。使用旧的PYTHONPATH时陷阱会触发。审阅注意事项:
dsv4-fp4-b200-sglang-agentic-hicache-mtp的 sweep 会在setup失败;这些 sweep 在运行时本来也会失败。glm5.2-fp8-mi325x-sglang-agentic-mtp和dsv4-fp4-mi355x-sglang-disagg-agentic-umbp-dspark在其配方修复之前,被 sweep 选中时会在setup失败。修正 dspark 的客户端镜像会改变 benchmark 实际运行的内容,因此需要单独的 changelog 记录和 sweep。validator 只检查 sweep 计划中的点,所以二者都不会挡住无关的 PR。single_node_environment复刻了benchmark-tmpl.yml把 matrix 字段导出给任务的方式;那边的映射改动时,这里也要同步修改。run-sweep只在修改perf-changelog.yaml的 PR 上触发,因此本 PR 自己的检查不会运行新的setup步骤,由上面的重放代替。perf-changelog.yaml记录,因为本 PR 不改任何 benchmark 配置或配方。@SemiAnalysisAI/core。AI 模型使用说明
关联 issue
无。#3567 的后续;相关 PR:#3428、#3446。
改动类型
新功能。